Goto

Collaborating Authors

 hugging face


OpenAI Gets Sued Over the Hugging Face Hack

WIRED

A nonprofit in California is doing what Hugging Face has not--attempting to hold OpenAI legally accountable for the actions of its agents. A legal nonprofit sued OpenAI in a California court on Tuesday over the company's agents escaping a testing environment and hacking the open source AI platform Hugging Face . "OpenAI's actions straightforwardly violated California law," the suit alleges. The suit was filed by Legal Advocates for Safe Science and Technology (LASST) and the law firm Gerstein Harrow in California Superior Court in San Francisco, where OpenAI is headquartered. It alleges that OpenAI's agents violated California's Comprehensive Computer Data Access and Fraud Act (CDAFA) by breaching Hugging Face over the summer.


Who's liable when AI agents go rogue?

MIT Technology Review

Who's liable when AI agents go rogue? Recent hacks have shown that the law is lagging when it comes to holding companies accountable. In July, OpenAI disclosed that a swarm of its agents had escaped their sandbox and hacked into the AI platform Hugging Face to cheat on a cybersecurity test. Recently, external researchers discovered that OpenAI agents had hijacked a German wiki site and the coding platform RubyGems in May to share test answers. Earlier this month, Anthropic disclosed four incidents in which its model Claude hacked into third-party systems during cybersecurity exercises. Just last week, Google confirmed that its model Gemini had been caught hacking other companies too.


AI agent kill switch urged by Okta-led alliance – how businesses could make it work

ZDNet

In response to the emerging threat of rogue and shadow AI agents, here's how the newly formed Blueprint Alliance seeks to help businesses secure their systems against anomalous AI activity. David Berlind is one of the founding editors of ZDNET and is an award-winning tech journalist. Okta, AWS, Google Cloud, Salesforce, and others form an AI agent security coalition. The Alliance offers a blueprint for companies seeking visibility, control, and governance of agents. AI agents need the equivalent of a kill switch to expeditiously terminate suspicious behavior.


OpenAI's agent hacked into an Australian government website

Engadget

An OpenAI agent hacked into the public website of the Australian government's Medicare public health insurance system in June, Australian Prime Minister Anthony Albanese has revealed. Albanese said he spoke with Sam Altman, the company's CEO, to express the country's "extreme concern" about the incident. He also criticized the company for taking too long to notify Australian authorities. According to The Guardian, OpenAI only sent an email about the breach to a general Australian government email address on September 10. And since authorities only check that account once a day, they didn't see it until September 11.


An OpenAI agent hacked the Australian government

Mashable

AI at School Mashable's Best: E-readers, robovacs, laptops, earbuds, smart home and more Mashable Selects Look Up Say More Safety Net Versus Creator Playbook In My Bag Trending Now Back to School Good Connection: Uplifting stories for a digital age All Series The AI model autonomously breached government systems to gain access to non-public information. Amanda Yeo is an Assistant Editor at Mashable, covering entertainment, culture, tech, science, and social good. Based in Australia, she writes about everything from video games and K-pop to movies and gadgets. An OpenAI agent has hacked the Australian government's healthcare system, gaining unauthorised access to non-public information as well as writing files to the server. SEE ALSO: Anthropic researcher quits, says AI'could kill us all by the end of the decade' Australian prime minister Anthony Albanese announced the breach at a press conference on Thursday .


OpenAI and Anthropic bosses push UN for global terms on AI

BBC News

Image caption, Sam Altman of OpenAI has become a key figure in the AI technology development. The heads of OpenAI, Anthropic, and Hugging Face have told the UN that the current pace of artificial intelligence (AI) development, and the risks it poses to society, demands international coordination. Altman called for common risk evaluation standards, as did Dario Amodei of Anthropic, a main rival of OpenAI, and Clement Delangue of Hugging Face. Earlier this month, Amodei wrote an essay welcomed by Altman and others calling an AI development slowdown in response to fears about the technology's threat to humanity. However, at the same UN conference, a key technology advisor to US President Donald Trump, rejected the idea any new form of AI regulation.


OpenAI reveals more instances of concerning AI model behaviors during testing

Engadget

OpenAI has revealed six incidents, wherein the models it was testing acted on their own and behaved in concerning ways it didn't expect, in a post about how it was adopting a new framework for "misalignment reports." In one one incident, the company said that a model found and used an exposed API key without permission while answering routine questions about earnings figures in a California county. When it still failed to find the figures, it fabricated them and presented them as facts from a legitimate source. If this had occurred in any other profession, we doubt the perpetrator would have much of a career for long. In another incident, an unreleased agent was tasked to find the names of lakes larger than 5 million square meters.


AI models chatting in 'surreal' dialect mixing poetic language and tech bro jargon

The Guardian

Interest in language used by AI models has increased since rogue OpenAI agents hacked into real-world company Hugging Face. Interest in language used by AI models has increased since rogue OpenAI agents hacked into real-world company Hugging Face. Experts air concern over AI lingo redolent of James Joyce's prose that is creating headaches for monitoring and oversight AI models have begun communicating in a strange new version of English that reads like a cross between James Joyce's Finnegans Wake and tech bro jargon, new research has found. Autonomous AI agents are rapidly creating novel dialects allowing them to converse in an often barely comprehensible language, which risks making it harder for humans to monitor their behaviour. Researchers at Emergence, a frontier AI lab in New York, found that within days of being asked to cooperate in experimental "societies", the models from several of the world's largest AI companies begin creating phrases, shorthands and agreed meanings they had never been explicitly taught.


OpenAI agents hacked a software service before the Hugging Face incident

Engadget

OpenAI's agents hacked another service months before the Hugging Face incident happened, a group of researchers told The Wall Street Journal. The agents, which the company was testing in a supposed sandbox environment, reportedly broke into RubyGems, which is a community-ran packaging service for Ruby programs and libraries. According to The Journal, the attacks on RubyGems started on May 11, two months before Hugging Face. The agents created accounts every two to three minutes and then uploaded hundreds of files to the service. RubyGems had to shut down account registration for four days in order to stop the attacks. Typically, creators on RubyGems upload files containing code and other information to help advance software development, but the agents' documents contained web pages scraped from the internet instead.


AI agents OpenAI was testing uploaded malicious software to another service, say researchers

The Guardian

Sam Altman speaks during a discussion with Howard Lutnick at a summit, in Chapel Hill, North Carolina, on 2 September 2026. Sam Altman speaks during a discussion with Howard Lutnick at a summit, in Chapel Hill, North Carolina, on 2 September 2026. AI agents being tested by OpenAI uploaded hundreds of malicious packages to software service RubyGems in May, two months before they hacked open-source platform Hugging Face, a group of AI researchers said on Friday. "On May 11th, 2026, hundreds of malicious packages were uploaded to RubyGems by AI agents. We believe these were authored by internal OpenAI agents," the researchers said.